Papers with monolingual tasks

7 papers
Recycling a Pre-trained BERT Encoder for Neural Machine Translation (D19-56)

Copied to clipboard

Challenge: In monolingual tasks, the number of unlearned model parameters is as huge as the number learned parameters in the BERT model.
Approach: They propose to apply a pre-trained Bidirectional Encoder Representations from Transformers (BERT) model to Transformer-based neural machine translation (NMT) based on the Transformer.
Outcome: The proposed model is stable and efficient in low-resource settings.
Mixed-Lingual Pre-training for Cross-lingual Summarization (2020.aacl-main)

Copied to clipboard

Challenge: Cross-lingual summarization (CLS) aims at producing a summary in the target language for an article in the source language.
Approach: They propose a mixed-lingual pre-training scheme that leverages both cross-lingual tasks such as translation and monolingual tasks like masked language models.
Outcome: The proposed model improves on the translation and masked language models with no task-specific components and saves memory.
Modular Sentence Encoders: Separating Language Specialization from Cross-Lingual Alignment (2025.acl-long)

Copied to clipboard

Challenge: Multilingual sentence encoders are often trained to map sentences from different languages into a shared semantic vector space . cross-lingual alignment training distorts optimal monolingual structure of semantic spaces of individual languages . a modular solution can be used for cross-linguistic tasks such as cross-language semantic similarity and zero-shot transfer .
Approach: They propose a modular training system that embeds sentences from different languages into a shared semantic vector space.
Outcome: The proposed solution achieves better performance across all tasks compared to monolithic models.
DLIR: Spherical Adaptation for Cross-Lingual Knowledge Transfer of Sociological Concepts Alignment (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for identifying nuanced sociological concepts fail to capture domain-specific subtleties or require extensive parallel data.
Approach: a new approach to aligning nuanced sociological concepts is proposed . a dual-branch LoRA approach captures core semantics and counteracts specific language perturbations.
Outcome: a new approach outperforms baselines on cross-lingual sociological concept retrieval across 10 languages.
KIT-Multi: A Translation-Oriented Multilingual Embedding Corpus (L18-1)

Copied to clipboard

Challenge: Cross-lingual word embeddings are representations of words across languages in a shared continuous vector space.
Approach: They propose a multilingual word embedding corpus which is acquired by neural machine translation and is based on monolingual data.
Outcome: The proposed method is competitive with existing methods but on the cross-lingual document classification task, it obtains the best figures.
LexCLiPR: Cross-Lingual Paragraph Retrieval from Legal Judgments (2025.acl-long)

Copied to clipboard

Challenge: Existing work on IR focus on retrieving entire cases rather than precise, paragraph-level information.
Approach: They propose a cross-lingual dataset for paragraph-level retrieval from ECtHR judgments . they evaluate retrieval models in a zero-shot setting and use multilingual case law guides .
Outcome: The proposed model excels in cross-lingual retrieval, while siamese architectures are better suited for monolingual tasks.
Multilingual Large Language Models Are Not (Yet) Code-Switchers (2023.emnlp-main)

Copied to clipboard

Challenge: Existing multilingual Large Language Models are not specifically trained with objectives for managing code-switching scenarios.
Approach: They propose to use multilingual Large Language Models to perform sentiment analysis, machine translation, summarization and word-level language identification to compare their performance to fine-tuned models of much smaller scales.
Outcome: The proposed models show that they underperform in comparison to fine-tuned models of much smaller scales.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations